Papers with automated methods
Predicting Depression in Screening Interviews from Latent Categorization of Interview Prompts (2020.acl-main)
Copied to clipboard
| Challenge: | Existing methods to diagnose depression require time-intensive interviews, assessments, and analysis. |
| Approach: | They propose a model that analyzes interview transcripts to identify depression while jointly categorizing interview prompts into latent categories. |
| Outcome: | The proposed model outperforms baseline models and provides psycholinguistic insights about depression. |
Knowledge Discovery and Hypothesis Generation from Online Patient Forums: A Research Proposal (P19-2)
Copied to clipboard
| Challenge: | Unprompted patient experiences on patient forums contain a wealth of unexploited knowledge. |
| Approach: | They propose to develop automated methods for mining, aggregating and cross-linking patient knowledge from online forums. |
| Outcome: | The proposed methods could be compared with biomedical literature and provide hypotheses for future clinical research. |
A Comparative Multidimensional Analysis of Empathetic Systems (2024.eacl-long)
Copied to clipboard
| Challenge: | Empathetic dialogue systems have received significant attention, but no systematic review has verified these limitations. |
| Approach: | They analyze 21 empathetic dialogue systems using automated methods to examine their progress. |
| Outcome: | The results show that empathetic dialogue systems lack specificity, reflection levels, diversity . the results also offer guidance for developing future systems . |
Deep Learning and Sociophonetics: Automatic Coding of Rhoticity Using Neural Networks (N19-3)
Copied to clipboard
| Challenge: | Automated extraction methods for vowels are available, but coding rhoticity has lagged behind. |
| Approach: | They use Neural Networks/Deep Learning to train a model on 208 speakers in Boston . they find that there is no reliable method for classifying r-dropping . |
| Outcome: | The proposed method trains a model on 208 speakers in Boston, Massachusetts. |
Concept Distillation from Strong to Weak Models via Hypotheses-to-Theories Prompting (2025.naacl-industry)
Copied to clipboard
Emmanuel Aboah Boateng, Cassiano O Becker, Nabiha Asghar, Kabir Walia, Ashwin Srinivasan, Ehi Nosakhare, Soundararajan Srinivasan, Victor Dibia
| Challenge: | Concept Distillation (CD) is an automated prompt optimization technique for enhancing weaker models on complex tasks. |
| Approach: | They propose an automatic prompt optimization technique for enhancing weaker models on complex tasks using a base prompt and a strong model to generate reasons for these mistakes. |
| Outcome: | The proposed technique improves weaker models on NL2Code and mathematical reasoning tasks, while preserving performance. |
DelucionQA: Detecting Hallucinations in Domain-specific Question Answering (2023.findings-emnlp)
Copied to clipboard
Mobashir Sadat, Zhengyu Zhou, Lukas Lange, Jun Araki, Arsalan Gundroo, Bingqing Wang, Rakesh Menon, Md Parvez, Zhe Feng
| Challenge: | Hallucination is a well-known phenomenon in text generated by large language models . state-of-the-art LLMs still have a number of weaknesses, including the tendency to generate hallucinatory statements without considering the factuality . |
| Approach: | They propose a dataset that captures hallucinations made by retrieval-augmented LLMs . they propose to use these methods to help detect hallucinosity in QA tasks . |
| Outcome: | The proposed method captures hallucinations made by retrieval-augmented LLMs for QA tasks. |
Dialogue-AMR: Abstract Meaning Representation for Dialogue (2020.lrec-1)
Copied to clipboard
Claire Bonial, Lucia Donatelli, Mitchell Abrams, Stephanie M. Lukin, Stephen Tratz, Matthew Marge, Ron Artstein, David Traum, Clare Voss
| Challenge: | Abstract Meaning Representation (AMR) does not capture the illocutionary force or speaker’s intended contribution in the broader dialogue context. |
| Approach: | They propose a schema that enriches Abstract Meaning Representation (AMR) it provides a semantic representation for facilitating Natural Language Understanding (NLU) in dialogue systems. |
| Outcome: | The proposed schema provides a semantic representation for facilitating Natural Language Understanding (NLU) in human-robot dialogue systems. |
Environmental Claim Detection (2023.acl-short)
Copied to clipboard
| Challenge: | a growing number of environmental claims are being made by companies in the face of climate change. |
| Approach: | They propose a task of environmental claim detection to detect environmental claims at scale . they use an expert-annotated dataset and models trained on this dataset to do this . |
| Outcome: | The proposed task detects environmental claims in quarterly earning calls . the number of environmental claims has steadily increased since the Paris Agreement in 2015 . |
Building Data-Driven Occupation Taxonomies: A Bottom-Up Multi-Stage Approach via Semantic Clustering and Multi-Agent Collaboration (2025.emnlp-industry)
Copied to clipboard
| Challenge: | Existing methods for creating robust occupation taxonomies are slow and expensive . a robust taxonomy is critical for job recommendation and labor market intelligence applications . |
| Approach: | They propose a framework that automates creation of occupation taxonomies from job postings . they use global semantic clustering to distill core occupations, then a reflection-based multi-agent system to iteratively build a coherent hierarchy. |
| Outcome: | The proposed framework produces taxonomies that capture unique regional characteristics. |
Contextualized Topic Coherence Metrics (2024.findings-eacl)
Copied to clipboard
| Challenge: | Existing topic models that estimate the interpretability of topics are difficult to compare due to their nature as unsupervised models. |
| Approach: | They propose to use contextualized topic coherence metrics to simulate human-centered coherency evaluation while maintaining the efficiency of other automated methods. |
| Outcome: | The proposed metrics better reflect human judgment on topics extracted from short text collections by avoiding highly scored topics that are meaningless to humans. |
Self-prompted Chain-of-Thought on Large Language Models for Open-domain Multi-hop Reasoning (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing open-domain question-answering methods lack quality assurance . existing methods lack scalability and poor diversity, hindering LLMs' capabilities . |
| Approach: | They propose an open-domain multi-hop reasoning framework to answer multi-choice questions . they propose an adaptive sampler for in-context selection and self-prompted inference . |
| Outcome: | The proposed framework surpasses the existing SOTA methods on large-scale datasets and doubles the zero-shot performance of small-scale LLMs. |
SurveyPilot: an Agentic Framework for Automated Human Opinion Collection from Social Media (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for opinion survey research exhibit severe biases and lack traceability. |
| Approach: | They propose a finite-state orchestrated agentic framework that automates the collection and analysis of human opinions from social media platforms. |
| Outcome: | The proposed framework achieves close alignment with authentic survey results across multiple domains, with average relative improvements of 68,98% and 51,37% when compared to opinion synthesis and agent-based approaches. |
Cross-lingual and Word-Independent Methods for Quantifying Degree of Grammaticalization (2026.eacl-long)
Copied to clipboard
Ryo Nagata, Daichi Mochihashi, Misato Ido, Yusuke Kubota, Naoki Otani, Yoshifumi Kawasaki, Hiroya Takamura
| Challenge: | Existing methods for quantifying the degree of grammaticalization are language- and word-dependent . existing methods are language dependent and lack training data . |
| Approach: | They propose to use Positive-Unlabeled learning or Cross-Validation-like learning to quantify degree of grammaticalization. |
| Outcome: | The proposed method achieves high correlations to human judgments in English deverbal prepositions and Japanese nouns being grammaticalized. |
A Top-down Graph-based Tool for Modeling Classical Semantic Maps: A Case Study of Supplementary Adverbs (2025.naacl-long)
Copied to clipboard
| Challenge: | Semantic map models (SMMs) construct a network-like conceptual space from cross-linguistic instances or forms based on the connectivity hypothesis. |
| Approach: | They propose a graph-based algorithm that automatically generates conceptual spaces and SMMs in a top-down manner. |
| Outcome: | The proposed model is compared with human annotations and other automated methods on cross-linguistic supplementary adverbs. |
Standardizing Distress Analysis: Emotion-Driven Distress Identification and Cause Extraction (DICE) in Multimodal Online Posts (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for identifying hate speech have been limited to analyzing textual content. |
| Approach: | They propose a method for distress identification and cause extraction from social media posts using emotional information. |
| Outcome: | The proposed method improves F1 and ROS scores by 1.95% and 3% relative to the best-performing baseline. |
CodeWiki: Evaluating AI’s Ability to Generate Holistic Documentation for Large-Scale Codebases (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing automated methods struggle to capture rich semantic dependencies and architectural structure. |
| Approach: | They propose a framework for automated repository-level documentation across seven programming languages. |
| Outcome: | The proposed framework outperforms the closed-source DeepWiki benchmark by 68.79% and is open source to support future research. |
Joint Turn and Dialogue level User Satisfaction Estimation on Multi-Domain Conversations (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to estimate turn and dialogue level user satisfaction employ hand-crafted features and rely on complex annotation schemes, which reduce generalizability of the trained models. |
| Approach: | They propose to use an adaptive multi-task loss function to minimize hand-crafted features to estimate user satisfaction at turn level from an end user perspective. |
| Outcome: | The proposed model improves on 28 Alexa domains, two dialogue systems and three user groups on a set of user-generated dialogues from 28 Alexia domain and 28 Alexis domains. |
Uncovering Latent Arguments in Social Media Messaging by Employing LLMs-in-the-Loop Strategy (2025.findings-naacl)
Copied to clipboard
| Challenge: | Supervised methods are adept at text categorization, but dynamic nature of social media debates pose challenges for them . traditional methods for extracting themes from public discourse often reveal overarching patterns that might not capture specific nuances. |
| Approach: | They propose a generic approach that leverages the advanced capabilities of Large Language Models to extract latent arguments from social media messaging. |
| Outcome: | The proposed approach leverages the advanced capabilities of Large Language Models (LLMs) to extract latent arguments from social media messaging. |
M-LongDoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework (2025.emnlp-main)
Copied to clipboard
Yew Ken Chia, Liying Cheng, Hou Pong Chan, Maojia Song, Chaoqun Liu, Mahani Aljunied, Soujanya Poria, Lidong Bing
| Challenge: | Existing benchmarks for large multimodal models focus on short documents with less than 50 pages and are limited to extraction-based questions. |
| Approach: | They propose a retrieval-aware tuning approach to improve the accuracy of multimodal document reading by 4.6%. |
| Outcome: | The proposed framework improves the accuracy of model responses by 4.6% compared to existing benchmarks on documents with hundreds of pages and longer documents with more complex content. |
A Lightweight Method to Generate Unanswerable Questions in English (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to build robust question answering models are too complex . antonym and entity swaps on answerable questions are used to build models . |
| Approach: | They propose a method for performing antonym and entity swaps on unanswerable questions. |
| Outcome: | The proposed method outperforms the previous state-of-the-art and has higher human-judged relatedness and readability. |
MiCEval: Unveiling Multimodal Chain of Thought’s Quality via Image Description and Reasoning Steps (2025.naacl-long)
Copied to clipboard
Xiongtao Zhou, Jie He, Lanyu Chen, Jingyu Li, Haojing Chen, Victor Gutierrez Basulto, Jeff Z. Pan, Hanjie Chen
| Challenge: | Existing methods for evaluating the quality of reasoning steps in multimodal chain-of-thought are lacking. |
| Approach: | They propose a framework to evaluate the correctness of reasoning chains by evaluating the quality of both the description and each reasoning step. |
| Outcome: | The proposed framework improves interpretability and human judgments on four state-of-the-art MLLMs. |
Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph Construction (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for knowledge graph creation (KGC) are limited in their ability to scale up to text common in many real-world applications. |
| Approach: | They propose a framework for knowledge graph creation from input text using a pre-defined schema and a trained component that retrieves schema elements relevant to the input text. |
| Outcome: | The proposed framework extract-define-canonicalize extracts high-quality triplets with a succinct self-generated schema without any parameter tuning and with significantly larger schemas compared to prior works. |
Generating Pedagogically Meaningful Visuals for Math Word Problems: A New Benchmark and Analysis of Text-to-Image Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Math word problems (MWPs) describe mathematical scenarios through text, requiring learners to interpret both linguistic and numerical information to derive mathematical expressions. |
| Approach: | They propose a framework for generating pedagogically meaningful visuals from MWP text descriptions using a pre-defined visual language and a design space grounded in interviews with math teachers. |
| Outcome: | The proposed framework illustrates the core mathematical relationships in math word problems. |
CodeAgent: Autonomous Communicative Agents for Code Review (2024.emnlp-main)
Copied to clipboard
Xunzhu Tang, Kisub Kim, Yewei Song, Cedric Lothritz, Bei Li, Saad Ezzini, Haoye Tian, Jacques Klein, Tegawendé Bissyandé
| Challenge: | Existing methods for code review rely on single input-output generative models and thus lack the collaborative nature of code review. |
| Approach: | They propose a multi-agent Large Language Model (LLM) system for code review automation that incorporates a supervisory agent to ensure that all the agents’ contributions address the initial review question. |
| Outcome: | The proposed system detects inconsistencies between code changes and commit messages, identify vulnerabilities, validates code style adherence, and suggests code revisions. |
People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection (2023.emnlp-main)
Copied to clipboard
| Challenge: | Past work has shown that counterfactually augmented data (CADs) can improve models' performance on out-of-domain tests. |
| Approach: | They use Polyjuice, ChatGPT, and Flan-T5 to automatically generate CADs . they find that CAD generates a model that flips the original label with minimal changes . |
| Outcome: | The proposed model improves model robustness on out-of-domain test sets and individual data points. |
Incorporating Zoning Information into Argument Mining from Biomedical Literature (2022.lrec-1)
Copied to clipboard
| Challenge: | Argumentative zoning is a text zonation scheme that is used to segment text into zones that serve distinct functions. |
| Approach: | They propose to use zoning information to incorporate into argument mining tasks . they add zonation labels predicted by an off-the-shelf model to the beginning of each sentence . |
| Outcome: | The proposed models improve argument mining models without additional annotation cost. |
IDEM: The IDioms with EMotions Dataset for Emotion Recognition (2024.lrec-main)
Copied to clipboard
Alexander Prochnow, Johannes E. Bendler, Caroline Lange, Foivos Ioannis Tzavellos, Bas Marco Göritzer, Marijn ten Thij, Riza Batista-Navarro
| Challenge: | idiomatic expressions are used in everyday language and typically convey affect, i.e., emotion. |
| Approach: | They present a dataset of idiom-containing sentences that were generated and labelled with any one of 36 emotion types using a generative language model. |
| Outcome: | The proposed method achieves an agreement rate of 62% on the IDioms with EMotions dataset, with human validation by two independent annotators. |
Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for audio captioning lack fine-grained detail and contextual accuracy due to limited unimodal or superficial information. |
| Approach: | They propose a two-stage automated pipeline that uses pretrained models to extract contextual cues from video . a large language model synthesizes these inputs to generate detailed and context-aware captions . |
| Outcome: | The proposed method is scalable and generates detailed and context-aware captions on large-scale audio datasets. |
SciImpact: A Multi-Dimensional, Multi-Field Benchmark for Scientific Impact Prediction (2026.findings-acl)
Copied to clipboard
| Challenge: | Prior work on scientific impact prediction has focused on citation counts and its variants, leaving limited evaluation of models’ capability to reason about other dimensions. |
| Approach: | They propose a large-scale, multi-dimensional benchmark for scientific impact prediction spanning 19 fields. |
| Outcome: | The proposed model outperforms larger models and close-source models in a wide range of fields and measures of scientific impact across 19 fields. |
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models (2026.acl-long)
Copied to clipboard
Yanbin Yin, Kun Zhou, Zhen Wang, Xiangdong Zhang, Yifei Shao, Shibo Hao, Yi Gu, Jieyuan Liu, Somanshu Singla, Tianyang Liu, Eric P. Xing, Zhengzhong Liu, Haojian Jin, Zhiting Hu
| Challenge: | closed-ended question-based benchmarks struggle with saturation as newer models emerge . crowd-sourced leaderboards rely on costly and slow human judges . |
| Approach: | They propose a framework that leverages collective intelligence from all large language models to evaluate each other. |
| Outcome: | a new framework enables a democratic, pairwise evaluation of all large language models . it achieves 97% correlation with human judgements, while significantly reducing the cost. |